Papers with abstract procedural
Multimodal Intent Discovery from Livestream Videos (2022.findings-naacl)
Copied to clipboard
Adyasha Maharana, Quan Tran, Franck Dernoncourt, Seunghyun Yoon, Trung Bui, Walter Chang, Mohit Bansal
| Challenge: | Existing models for instructional video understanding struggle to understand abstract intents . identifying procedural intent within instructional videos is a challenging task . |
| Approach: | They propose to extract instructional intent from software instructional livestreams by using a multimodal cascaded cross-attention model that integrates weaker and noisier video signals with more discriminative text signals. |
| Outcome: | The proposed model improves on baseline models and compares it to existing models. |
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Captioning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation metrics fail to evaluate factual correctness in procedural video captions . Existing metrics rely on lexical overlap or holistic semantic similarity, but miss role-specific omissions resulting in hallucinations . |
| Approach: | They propose a role-aware, fact-level evaluation framework that distinguishes conceptual facts from contextual facts. |
| Outcome: | Experiments show that state-of-the-art captioning models produce fluent but incomplete descriptions with systematic errors. |